Fleet changelogs · dev.ecs0.net
jdmbair13m5 - incident - container-staging-eloop-total-app-launch-outage - claude - system-triage - fleet - 20260822-2320

jdmbair13m5 — total app-launch outage (container staging ELOOP)

When: diagnosed 2026-08-22 23:06–23:22 EDT · Host: jdmbair13m5 · Agent: claude (Opus 5) Verdict: NOT malware. macOS sandbox-container path bug from a half-completed account rename.

One-line summary

Every sandboxed app launch failed and Finder deadlocked because containermanagerd was being handed the home path /Users/rich (a compatibility symlink) instead of the real /Users/richh, and it refuses to create a sandbox staging directory through a symlink — returning errno 62 ELOOP.

Evidence chain

  1. sample of Finder (pid 4596): main thread 3884/3884 samples blocked in -[NSWindow _doOrderWindow:] → -[NSPersistentUIManager hasPersistentStateToRestore] → -[NSPersistentUIRemoteStorageClient stateDirectoryAtLaunch] → dispatch_sync → kevent_id.
  2. The queue it waits on (NSPersistentUIRemoteStorageClient) is itself parked in __NSXPCCONNECTION_IS_WAITING_FOR_A_SYNCHRONOUS_REPLY__ → mach_msg2_trap — a sync XPC call that never returns. Classic two-hop deadlock; the reply never comes because the container query underneath it fails.
  3. Unified log, continuous and ongoing at diagnosis time: CREATE_STAGING_DIRECTORY at path [/Users/rich/Library/ContainerManager/Staging/<uuid>] with errno (62) Too many levels of symbolic links emitted by secinitd on behalf of every launching agent (SSMenuAgent, recentsd, contactsdonationagent, WallpaperAgent, NotificationCenter, syncdefaultsd, WiFiAgent …), with PIDs climbing continuously — launchd crash-looping them.
  4. libsystem_secinit cannot initialise the sandbox → the process is killed at launch. That is the "all apps crash" symptom. Finder's hang is the same failure, one layer up.

Why the path was wrong — and why it is NOT persistent

Every on-disk record is already correct:

Source Value
dscl . -read /Users/richh NFSHomeDirectory /Users/richh
getpwuid(502).pw_dir /Users/richh
/etc/passwd no rich* entry
launchctl getenv HOME unset
grep /Users/rich in /Library/Launch*, ~/Library/LaunchAgents, /etc, ~/Library/Preferences, shell rc no hits

/Users/rich exists only as a single, non-looping symlink to /Users/richh (created 2026-08-21 06:57). It resolves correctly for ordinary processes — verified by creating and removing a probe directory through it. containermanagerd rejects it by policy, not because the filesystem loops: it will not create a container staging directory through a symlinked path.

The stale /Users/rich therefore lived only in the running login session — secinitd pid 75935 (started 2026-08-21 ~22:47, 26 minutes of CPU burned retrying). A reboot is the correct and sufficient remedy; nothing on disk needed editing.

Lighter-weight alternative for next time, not used here because a reboot was requested: sudo killall secinitd — launchd respawns it and it re-reads the corrected home path.

Security sweep — CLEAN

Owner action items

  1. Hostname drift — this host reports jdmbair13m5; fleet docs list rdmbair13m5. Correlates with the rich → richh rename. Decide which is intended and reconcile fleet docs.
  2. Remote Apple Events (eppc, tcp/3031) is enabled. Rarely needed; consider turning it off.
  3. ~/Library/Saved Application State is missing (collateral); macOS recreates it on next launch.
  4. Apple Notes entry to folder llmlog is PENDING — notes_changelog.zsh refuses to run from an SSH session. File it from Terminal.app in the desktop session on this host.

Fleet sync performed in the same session

File Before After Source
~/.claude/CLAUDE.md 37,777 B 48,503 B rdmsm4x (strict superset — no local headings lost)
~/dev_update.zsh 66,420 B 68,365 B rdmsm4x (higher version wins)
~/.claude/settings.json verbose unset verbose: true fleet baseline key that had drifted

AGENT_COORDINATION.md and FLEET.md were already byte-identical to rdmsm4x.

Deliberately NOT synced: per-host settings keys (editorMode, enabledPlugins, hooks, skillOverrides, statusLine, theme, worktree, permissions.defaultMode) and rdmsm4x's 48 permissions.allow rules — widening what runs without asking is Rich's decision, not a sync side effect.

Backups / undo: ~/.claude/backups-20260822-2311/{CLAUDE.md,dev_update.zsh,settings.json}.pre-sync Restore with cp ~/.claude/backups-20260822-2311/CLAUDE.md.pre-sync ~/.claude/CLAUDE.md etc.

Raw evidence: ~/Library/Logs/claude-incident-20260822-2311/ (finder-sample.txt, secinitd-sample.txt, cmd-sample.txt, secd-sample.txt, log-errors.txt.gz)

Check-in: ~/.agent-coordination/checkins/claude-jdmbair13m5-eloop-container-outage-20260822-2320.json (also published to rdmsm4x)